Back

Artificial Intelligence in the Life Sciences

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match Artificial Intelligence in the Life Sciences's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Making Accelerating Medicines Partnership Data Findable and Interoperable through a Common Data Model: Extending OMOP for Multi-Source Multimodal Data

Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.

2026-09-02 genetic and genomic medicine 10.64898/2026.08.31.26361831 medRxiv
Top 0.1%
7.8%
Show abstract

SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.

2
Back to basics: Observed statistics are sufficient to predict drug responses

Svensson, V.; Khan, U.; Heydari, H.; Ubas, A. A.; Thomas, N.; Merico, D.; Goodarzi, H.; Yu, J.; Alidoust, N.; Gandhi, S.

2026-06-12 genomics 10.64898/2026.06.09.731197 medRxiv
Top 0.1%
6.2%
Show abstract

Predicting how cells, tissues, and patients will respond to a drug, cytokine, or genetic perturbation is central to biological and clinical reasoning. The practical goal is to estimate the analysis-ready readouts that support this reasoning: which cellular responses are context-dependent, which perturbations reveal shared or divergent mechanisms, and which observations should become the basis for the next experiment. Here we introduce Rhaister, a perturbation-response predictor that operates directly on screen-level summary statistics. By measuring just a few perturbations in a new biological context, Rhaister predicts the unmeasured perturbations by learning how response patterns vary across reference contexts. This formulation applies to both fine-grained molecular readouts, such as transcriptional responses from Tahoe-100M or other large perturbation screen, or on phenotypic endpoints. To train and apply Rhaister on pheno-typic endpoints we created Emerald Bay, a purpose-built Tahoe dataset that unifies multi-day cancer drug perturbation, pooled Mosaic tumor contexts, and paired transcriptomic response measurements*. Across these settings, Rhaister matches or exceeds substantially more expensive virtual-cell models, often achieving the highest values possible in evaluation metrics, while training in seconds and running predictions in milliseconds. On Emerald Bay, Rhaister predicts context-specific drug phenotypes from sensitivity measurements alone and improves further when including transcriptomic information. We also introduce Rhaister-O, predicting drug responses in new contexts from baseline expression alone and, to our knowledge, provides the first zero-shot model for this task. Rhaister establishes summary-statistic perturbation modeling as a fast, interpretable framework for predicting biological response across new contexts.

3
MorphoStat: A Statistics-Aware Pipeline for Morphological Profiling Analysis

Altobi, A.; Heo, D.

2026-06-18 bioinformatics 10.64898/2026.06.15.732111 medRxiv
Top 0.1%
4.1%
Show abstract

High-content imaging produces thousands of morphological measurements per cell. Interpreting these measurements requires normalization to remove plate effects, statistical tests selected on the basis of data distribution, and control over false discoveries across many features tested at once. MorphoStat is an open-source Python pipeline that applies this sequence of steps automatically. Given a CSV file from CellProfiler or a compatible imaging platform, it removes low-quality wells, normalizes each plate against DMSO controls using a MAD-scaled z-score, routes each feature to a parametric or nonparametric test based on a distributional check, applies Benjamini-Hochberg correction, and writes out results and publication-ready figures. On the BBBC021 benchmark (MCF-7 breast-cancer cells, 632 wells, 473 features), MorphoStat recovered 12 of 13 known mechanism-of-action classes in principal component space, confirming that the normalization and statistical routing work as intended. The tool is available at https://github.com/Almunthir334/morphostat (DOI: 10.5281/zenodo.20354069) under the MIT license.

4
Identifying and Addressing Systematic Data Leakage in Protein-Ligand Affinity Benchmarks

Mattsson, B.;Walters, W.

2026-06-30 Molecular Biology 10.64898/2026.06.29.735309 medRxiv
Top 0.1%
3.3%
Show abstract

Accurate prediction of protein-ligand binding affinity is a crucial goal in structure-based drug discovery, with the potential to significantly shorten development timelines. Recently, a new wave of machine learning models based on co-folding, such as Boltz-2 and IsoDDE, has demonstrated performance that matches or exceeds that of gold-standard physics-based methods like Free Energy Perturbation (FEP). This paper provides a critical assessment of these claims, revealing that current benchmarks are heavily influenced by data leakage, and proposes a new benchmark that explicitly controls for data leakage. We demonstrate that splitting by protein-sequence identity is inherently insufficient to prevent data leakage due to "target mirroring," in which homologous proteins with low overall sequence identity still exhibit highly correlated binding profiles. Our meta-analysis of documents in the ChEMBL 36 database identifies more than 6,000 such assay pairs and finds that leakage persists for sequence-identity thresholds as low as 0.2, well below the values commonly used in benchmarks today. Additionally, we show that a ligand-only baseline model, which lacks protein structural information, achieves surprisingly high performance on the FEP+ 4 and OpenFE benchmarks (r = 0.66 and r = 0.36, respectively). Our results indicate that current benchmarks tend to reward models for memorizing training data and exploiting localized leakage rather than truly learning biophysical principles. To address this issue, we propose the Novelty-Tiered Affinity Benchmark, in which the test data is partitioned into ligand novelty tiers. In the most challenging tier (Tanimoto similarity < 0.35), ligand-only models perform notably worse (r = 0.14), offering a clear baseline for evaluating genuine generalization. We argue that the field must move beyond sequence-based splits to ensure that AI-driven discovery translates into successful prospective laboratory research.

5
CHIMIYA-1: An Autoselection Foundation Model for ADMET Property Prediction, Rigorously Benchmarked Against the Therapeutics Data Commons ADMET Group

Varghese, R.; Tiwary, P.; Oswal, K.

2026-07-23 bioinformatics 10.64898/2026.07.20.739289 medRxiv
Top 0.1%
3.2%
Show abstract

Accurate, generalizable prediction of absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties remains one of the highest-leverage unsolved problems in computational drug discovery, and late-stage attrition driven by ADMET liabilities continues to be a dominant cost driver in pharmaceutical research and development. The Therapeutics Data Commons (TDC) ADMET Group has emerged as the fields most widely adopted public benchmark, comprising 22 endpoints under standardized scaffold-split evaluation. In this work we report a comprehensive evaluation of CHIMIYA-1, a proprietary autoselection foundation model developed by Covenant Biosciences, against the full TDC ADMET Group. Departing from common practice in the field, every reported score is the mean and standard deviation of five independently seeded end-to-end evaluation runs (TDCs own minimum submission standard, which we find is not met by all public leaderboard entries), and all 22 endpoints were additionally subjected to an explicit train/test structural-overlap audit prior to reporting, finding zero overlaps on any endpoint. Despite this deliberately conservative evaluation standard, CHIMIYA-1 ranks first among all publicly listed methods on four endpoints, places within the top decile of the field on twenty of twenty-two endpoints (91%), and attains a mean percentile standing near the 74th percentile across the full benchmark, with particular strength on toxicity and physicochemical-property endpoints. We further show that several top-ranked public comparators on this benchmark have been independently found to exhibit confirmed data leakage, a finding that, if anything, understates CHIMIYA-1s relative standing. All results were obtained on commodity single-GPU workstation hardware without recourse to distributed or cloud-scale training infrastructure. We discuss these results in the context of benchmark reporting norms in molecular machine learning and outline ongoing extensions, including continuous prospective-data retraining and CUDA-level throughput optimization of the underlying selection pipeline.

6
Real Science Is Harder Than Benchmarks: Evaluating Advanced AI Frameworks on Published Studies. I. Uncertainty Quantification, ML on Therapeutic Data Commons, and Agent-Based Modeling

Ahmed, M. O.; Amale, S. A.; Bhavsar, R. D.; Chopra, P.; Jaimes, A.; Kachhwah, A.; Kalotra, C. D.; Li, P.; Li, X.; Liao, Y.; Roy, R.; Senthilselvan, N.; Shao, Y.; Sharma, A. D.; Shrivatsan, A.; Xue, R.; You, Y.; Badkul, A.; Xie, L.; Oet, M.; Lee, K.; Sinitskiy, A.

2026-06-27 bioinformatics 10.64898/2026.06.24.734302 medRxiv
Top 0.1%
2.8%
Show abstract

Artificial Intelligence (AI) frameworks for automating scientific research have shown strong performance on benchmarks, but their capacity to routinely reproduce results from multiple real-life published studies remains largely untested. We evaluated five advanced AI research frameworks (Kosmos, K-Dense, ToolUniverse, BioAgents from bio.xyz, and the AI Scientist-v2 from Sakana AI) on three real-life tasks (including two recently published papers) spanning uncertainty quantification for molecular property predictions, machine learning on Therapeutic Data Commons benchmarks, and agent-based modeling. AI frameworks demonstrated genuine strengths: generating original hypotheses, competently executing routine data acquisition and coding tasks, providing statistical measures of confidence often absent from the original papers, and producing well-formatted final reports. At the same time, our experiments revealed that real-world scientific tasks remain considerably harder than current benchmarks suggest. No AI framework matched the scope or depth of the original studies, results varied across multiple runs of the same framework with the same prompt, and we documented cases of severe hallucinations in final reports, gaps in literature coverage, and overconfident conclusions. Verification of AI outputs required substantial domain expertise. While these three tasks are only partially representative of the broader scientific landscape, they offer a starting point for developing a more rigorous methodology for evaluation of AI performance than what is currently practiced. We conclude that AI frameworks are already valuable for prototyping research directions and stress-testing completed studies, and some of the limitations documented here appear largely tractable through infrastructure improvements and continued development.

7
Practical Use of Advanced AI Frameworks on Real-Life Scientific Problems: Three Case Studies

Gulluoglu, H. S. A.; Baby, J.; Bagul, K. M.; Basangari, B. R.; Bathini, S. A.; Chalamalla, N. K. R.; Dcunha, J.; Gupta, O.; Huang, L.; Jiang, X.; Naidu, Y. R.; Sathishkumar, G.; Sehrawat, M.; Thota, S. L.; Thuvara, D.; Vanguri, M. B.; Yin, J.; Jugder, B.-E.; Lusky, I. E.; Li, J.; Sinitskiy, A.

2026-06-29 bioinformatics 10.64898/2026.06.23.734132 medRxiv
Top 0.1%
2.7%
Show abstract

Agentic artificial intelligence (AI) systems increasingly claim to automate scientific research, yet independent evaluations report persistent gaps between those claims and demonstrated capability. We tested frontier agentic AI systems on three practical problems: prediction of treatment non-response in immune-mediated inflammatory diseases, optical chemical structure recognition for literature mining, and prediction of drug-design-related properties from small datasets. Each problem was first assigned to autonomous frameworks and then reattempted as human-led, AI-assisted work. Autonomous runs failed in most cases, while human-led work produced reusable resources and modest but defensible performance, including new evidence for possible mechanisms of treatment resistance and a more practical benchmark for mining chemical structures from scientific papers. Property prediction was the single task on which one autonomous AI framework matched the human expert. We conclude that current frameworks can carry out engineering and analysis once a human expert leads the project, but cannot yet engineer a novel solution without oversight. The use of AI on real-life scientific problems remains an art rather than a routine technology.

8
Systematic AI-Driven Drug Repurposing via Clinical Trial Data Mining: A Framework and Six Cross-Therapeutic Case Studies.

Gote, V.

2026-06-14 bioinformatics 10.64898/2026.06.11.731629 medRxiv
Top 0.1%
2.1%
Show abstract

Drug repurposing -- the application of approved or shelved compounds to new therapeutic indications -- offers a cost- and time-efficient alternative to de novo drug discovery. However, the systematic identification of repurposing candidates from the rapidly expanding body of clinical trial data remains a significant challenge. Here I present a publicly accessible AI-powered tool that mines the ClinicalTrials.gov registry to identify approved drugs with under-explored therapeutic potential in high-value disease areas. The tool integrates natural language processing, mechanism-of-action pathway analysis, and trial density scoring to surface candidates where biological plausibility is high and clinical trial coverage is sparse. I demonstrate the tools utility across six cross-therapeutic case studies spanning oncology, cardiology, neurology, rare diseases, immunology, and infectious disease. Key findings include: the identification of Zonisamide as an under-explored combination candidate for obesity alongside GLP-1 receptor agonists; mechanistic validation of SGLT2 inhibitors in heart failure with preserved ejection fraction (HFpEF); and a novel cross-domain mapping of anti-TNF biologics to early-stage neurodegeneration via shared neuroinflammatory pathways. The tool is freely accessible and designed to lower the barrier for academic and industry researchers to systematically pursue repurposing opportunities.

9
Residual Multi-Modal Learning for Pan-Breast-Cancer Drug Response Prediction

Huang, B.; Tasaka, L.; Li, J.; Islam, T.; Zhang, S.

2026-07-08 bioinformatics 10.64898/2026.07.03.736239 medRxiv
Top 0.1%
1.9%
Show abstract

Predicting drug sensitivity across diverse cancer cell lines remains a fundamental challenge in precision oncology, particularly for data-scarce cell lines where per-cell-line models overfit and lookup-table approaches cannot generalise to unseen biological contexts. We present DL4DR, a Two Tower Residual Late Fusion deep learning model that addresses this challenge through content-based, identity-free genomic conditioning. The Cell Line Tower encodes each cell line as a 3 x 139 x 139 genomic image - encoding gene expression, mutation severity, and copy-number variation as RGB channels - using a convolutional encoder that maps directly from biological content, never from a cell line ID. The Compound Tower combines three complementary molecular representations: D-MPNN graph message passing, ORNN octave convolutional image features, and an ECFP hard-memorization head that preserves activity-cliff resolution. Predictions are composed as a residual sum: f = fhard + {lambda}(zc). fresidual, where the learned gate $\lambda$ modulates how much interaction signal supplements the memorization baseline. Evaluated across 51 breast cancer cell lines(136,342 records), Residual Fusion outperforms the ECFP-Only baseline in 48/51 cell lines (94.1%), with {Delta}R2 > 0.02 in 26/51 (51.0%). On the leave-cell-line-out split - the decisive test of genomic generalisation - the mean {Delta} R2 = 0.016 across all 51 lines demonstrates that the genomic encoder learns transferable biological signal beyond cell line identity. External validation on 601 cell lines across 27 cancer tissue types (CellTiter-Glo dataset; 0 cell line overlap with training) achieves median R2 = 0.627, within the range of the internal random-split performance (R2 = 0.61--0.69), confirming pan-cancer generalisation. GradCAM interpretability on the Cell Line Tower recovers TP53 among the top-five cross-cell-line genomic activators (5/51 cell lines) alongside several uncharacterised candidate genes (e.g.FSIP2, 6/51) - without any prior pathway annotation - providing partial biological validation of the learned representation, while also indicating that a substantial share of the encoder's top-ranked signal corresponds to genes with no current annotation as breast cancer drivers. Code and data are available at https://github.com/bayjuan5/DL4DR.

10
From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology

qin, y.; Pang, J.; Zhang, X.

2026-09-01 bioinformatics 10.64898/2026.08.26.747436 medRxiv
Top 0.1%
1.9%
Show abstract

Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.

11
Machine Learning-Guided Discovery of Bacterial-Selective Membrane-Active Compounds Reveals Mechanistic Bias in Antibiotic Training Datasets

Chain, C.; Ghaffari, S.; Belakaria, S.; Sheehan, J. P.; Irani, I.; Wu, C.-Y.; Kim, H.; Engelhardt, B. E.; Gitai, Z. E.

2026-06-11 bioinformatics 10.64898/2026.06.08.730938 medRxiv
Top 0.1%
1.9%
Show abstract

The rise of antibiotic resistance necessitates the discovery of antibacterial compounds with novel mechanisms of action (MoAs). Recent machine learning approaches have shown promise in antibacterial compound discovery, but often identify derivatives of known antibiotic classes rather than mechanistically novel compounds. Previous approaches applied Tanimoto similarity filters at the end of screening pipelines, but this method has substantial drawbacks: Tanimoto similarity can be misleading in chemical space, and post-hoc filtering does not influence what activity models learn to prioritize. Here, we present a machine learning pipeline that addresses chemical novelty upfront by employing an XGBoost-based MoA classifier to explicitly prioritize compounds predicted to have mechanisms distinct from known antibiotic classes, combined with graph neural networks for antibacterial activity and toxicity prediction. Applied to the Zinc20 database, our approach successfully identified non-toxic antibacterial compounds structurally distinct from known antibiotics. Notably, the majority of these hits exhibited membrane-targeting activity with selectivity for bacterial cells over mammalian cells, suggesting potential for next-generation membrane-active antibiotics. However, we did not identify compounds with novel protein targets. Systematic analysis revealed that this limitation stems from mechanistic bias in training data rather than model architecture. Specifically, our activity model learned to preferentially score compounds similar to specific groups in the training data, thus overrepresenting certain MoA classes including membrane-active compounds. Even substantial model architecture and training data enhancements did not overcome this constraint. Our findings demonstrate that the primary bottleneck for discovering mechanistically novel antibiotics is the scarcity of diverse, mechanistically-annotated training data. This work provides both a methodological framework for mechanism-aware screening and critical insights into data requirements for genuinely novel antibiotic discovery.

12
Retention, not flux: endpoint confounding caps computational prediction of peptide skin penetration, with a delivery-aware reframing

Komianos, N.; Prakash, P.

2026-06-29 bioinformatics 10.64898/2026.06.25.734657 medRxiv
Top 0.1%
1.8%
Show abstract

Bioactive peptides are now central to cosmetic and dermatological actives, yet predicting whether a given sequence will reach its site of action in skin remains unsolved. We contend that the dominant framing, predicting a single binary "skin permeability" label from sequence, is ill-posed, and that this, rather than a shortage of modelling power, explains the field's stalled predictive performance. The scope of the claim is narrow: barrier-crossing propensity is a legitimate, learnable function of molecular structure, whereas the vehicle- and endpoint-agnostic binary label that the literature supplies is not. We support this with a first-principles analysis and a study of public-source data. First, the experimental endpoint most commonly reported, transdermal flux into a diffusion-cell receptor compartment (OECD Test Guideline 428), conflates two opposite outcomes (genuine deep delivery and undesired systemic transport) and is, for a cosmetic active, frequently a failure signal rather than a success signal. That receptor flux is an imperfect measure of cutaneous bioavailability is long established in dermatopharmacokinetics; our contribution is to show that the same confound, inherited through scraped labels, is what caps machine learning from sequence. Second, reported "permeability" is a property of the sequence x delivery-vehicle x measurement-compartment triad, two terms of which are usually unrecorded. Third, on public-source data, a physicochemical intrinsic-permeability estimate (Potts-Guy) carries no positive predictive signal for scraped penetration labels (grouped AUC 0.45, 95% CI 0.40-0.51); sequence-only classifiers plateau in the mid-0.70s with diminishing returns as labels accumulate (AUC 0.70-0.77); and the same descriptor pipeline on a clean single-endpoint membrane dataset scores materially higher (AUC 0.83, non-overlapping CI). Our proposed reframing separates barrier-crossing (data-driven, sequence-level) from depth-and-retention (physics-driven, delivery-aware) and treats intrinsic transdermal flux as a regulatory risk axis; we close by proposing a triad-annotated reporting schema and a seed benchmark.

13
CARD:Epi - Contextualizing Antimicrobial Resistance Determinants Using Deep Learning Language Models

Edalatmand, A.; Ta, T. E.; Zhao, C.; Ibrahim, A.; Upadhyaya, R.; Rajapaksa, S.; Raphenya, A. R.; McArthur, A. G.

2026-08-18 genomics 10.64898/2026.08.14.744850 medRxiv
Top 0.1%
1.5%
Show abstract

Bacterial outbreak publications outline the key factors involved in the uncontrolled spread of infection. Such factors include the environment, pathogens, hosts, and antimicrobial resistance genes (ARGs). Individually, each paper published in this area gives a glimpse into the devastating impact drug resistant infections have on healthcare, agriculture, and livestock. When examined together, these publications provide contextual information on ARG transmission, from the discovery of new resistance genes to their dissemination to different pathogens, hosts, and environments. We have extracted this information from publications in PubMed by using the biomedical deep-learning language model, BioBERT. We trained BioBERT on two tasks: entity recognition to identify AMR-relevant terms (i.e., ARGs, taxonomy, environments, geographical locations, etc.) and relation extraction to determine which terms identified through entity recognition contextualize ARGs. By collating results from 204,094 antimicrobial resistance publications worldwide, we have generated interpretable results about the sources where genes are commonly found. To visualize the dataset, we have created two pipelines to analyze transmission patterns of ARGs across agriculture, environments, and human populations using a Confusogram and Uniform Manifold Approximation and Projection. Overall, we have taken a large-scale approach to collect antimicrobial resistance data from a commonly overlooked resource, i.e., the systematic examination of the large body of AMR literature and have visualized how scientific literature can be used to assess transmission patterns of ARGs.

14
A Graph-based QSAR Modeling Pipeline for Predicting In vitro PubChem Assays and In vivo Human Hepatotoxicity: Mechanistic Analysis of Caspase-3/7 Activation

Chitikela, Y.; Zhu, c.; Jia, Z.

2026-06-12 bioinformatics 10.64898/2026.06.10.731399 medRxiv
Top 0.1%
1.5%
Show abstract

BackgroundCaspase-3 and -7 are key effector caspases in the apoptotic pathway, a form of programmed cell death, and their activities serve as a well-established biomarker for evaluating environmental chemical toxicity and informing chemical risk assessment. Loss of mitochondrial membrane potential is a key event in the activation of Caspase-3/7 signaling and the subsequent induction of apoptosis. Therefore, simultaneous assessment of mitochondrial membrane potential and Caspase-3/7 activity enables elucidation of the mechanisms and pathways through which apoptosis is initiated.. Rapid and accurate assessment of the potential toxicity of environmental chemicals and drugs remains a major challenge. Quantitative Structure-Activity Relationship (QSAR) modeling have been widely used for toxicity prediction. Graph-based approaches encode compounds directly as molecular graphs, allowing structure-activity relationships to be learnt from molecular topology without the information loss in binary fingerprints. While advanced graph models such as graph transformers (GTs) have shown outstanding performance in many domains, they have not been fully leveraged in QSAR modeling on Caspase and mitochondrial toxicity. MethodsWe propose a QSAR modeling pipeline that encompasses assay data preprocessing, feature representations (fingerprints and molecular graphs), and benchmarking machine learning (ML) models, including classic ML models, graph neural networks (GNNs), GTs, and their consensus ensembles. Based on in vitro Caspase and mitochondrial assays in PubChem, we applied the pipeline to predict Caspase-3/7 activation and mitochondrial membrane potential (MMP). Beyond in vitro assays, we also built in vivo QSAR modeling for FDA Drug-Induced Liver Injury (DILI) gold standard on human hepatotoxicity. Moreover, mechanistic analysis on Caspase-3/7 activation was conducted by comparing with MMP disruption to identify chemical substructures that may be responsible for dual activations. We also investigated cell-line-specific responses by identifying structural motifs that selectively induce Caspase-3/7 activation in individual cell lines. ResultsExperimental evaluations show that GTs and GNNs outperformed classic ML models when the number of active compounds is large, such as MMP disruption, while classic ML models and GTs performed good for highly imbalance data with limited active compounds, such as Caspase-3/7 activation. For DILI prediction, the full consensus model achieved the highest AUC 0.69 and Graphormer had the highest F1 score 0.79, both surpassing the previous best model with AUC 0.63 and F1 0.65 with a large margin. Our mechanistic analysis shows that phenolic compounds bearing a para-hydroxyphenyl motif, as well as members of the lipophilic chain family with long alkyl chains can trigger the collapse of MMP, leading to the activation of caspases-3 and -7. Human embryonic kidney (HEK293) was the only cell line with a distinct structural motif: 1,1-dichloroethane and chlorobenzene. Human neuroblastoma (SK-N-SH) is uniquely impacted by an epoxide fragment and rat hepatoma (H-4-II-E) is uniquely impacted by a tetramethylcyclohexene motif and an acetaldehyde fragment. ConclusionsThe proposed pipeline for QSAR modeling, including data preprocessing, feature representations, and incorporation of advanced graph ML approaches, is highly effective in predicting not only on Caspase-3/7 activation and membrane potential collapse, but also on FDA DILI human hetatotoxicity. As future research directions, we will leverage extra information, e.g., biological activity and findings in existing toxicity literature, and recent advances in large language models and agentic AI to further improve the predictive performance and enable a sensitive and specific framework for assessing human hepatotoxicity of environmental compounds.

15
B-SMART-Former: An Explainable Transformer-Based Deep Learning Model for Predicting Drug-Drug Interactions Between Biotech and Small-Molecule Drugs

Nasiri, F.; Hooshmand, M.; Nouroozi, M.

2026-07-27 bioinformatics 10.64898/2026.07.23.740240 medRxiv
Top 0.1%
1.5%
Show abstract

1Drug--drug interactions between biotech and small-molecule drugs play a critical role in medication safety and therapeutic efficacy. However, most existing computational DDI prediction methods focus primarily on interactions between small-molecule drugs, leaving biotech-small-molecule interactions comparatively underexplored. In this study, we propose B-SMART-Former, an explainable deep learning framework for predicting interaction types between biotech and small-molecule drugs. The proposed framework integrates ChemBERTa embeddings and Morgan molecular fingerprints for small molecules with ProtBERT embeddings for biotech drugs, eliminating the need for similarity-based features while leveraging complementary molecular representations. These multimodal features are processed by a hybrid architecture that combines Transformer-based self-attention, residual convolutional learning, and a multi-layer perceptron classifier to capture both global contextual dependencies and local discriminative patterns. The model is formulated as a multi-class classification task and evaluated using stratified 10-fold cross-validation. To improve model transparency, Integrated Gradients is employed as a post-hoc explainability method to identify the molecular features that contribute most strongly to each prediction. Experimental results demonstrate that B-SMART-Former achieves a micro-averaged AUROC of 0.9978 and an AUPR of 0.9682 while relying solely on intrinsic molecular representations, remaining competitive with similarity-based approaches. The proposed framework offers an effective and explainable solution for biotech-small-molecule DDI prediction and provides a practical foundation for future computational drug interaction studies.

16
PandaDock: An Open-Source Molecular Docking Platform with Flexible-Ligand Search and Equivariant Neural Scoring

Panda, P. K.

2026-08-20 bioinformatics 10.64898/2026.08.19.745667 medRxiv
Top 0.1%
1.4%
Show abstract

We present PandaDock, an open-source molecular docking platform implementing flexible-ligand conformational search with analytic gradients, a precomputed affinity grid engine, specialized modules for induced-fit, metal-coordination and tethered docking, and an SE(3)-equivariant graph neural network scoring function trained at scale. Ligand flexibility is represented as a torsion tree and pose parameters are optimized by Monte Carlo with Metropolis acceptance refined by L-BFGS, with rotational gradients obtained in closed form through the derivative of the SO(3) exponential map rather than by finite differences. Affinity grids are built by a blocked neighbor-selection scheme that is exact and 5.6-9.7x faster than dense evaluation, and may be cached across ligands sharing a receptor and site, reducing a six-ligand series from 29.3 s to 10.4 s. On 814 protein-ligand complexes spanning 14 target families, PandaDock recovers a pose within 2 Angstroms of the crystal geometry in 33.7% of cases at rank 1 and in 57.0% of cases within the returned ensemble. The GNN scoring function is trained on 741,706 co-folded complexes from SAIR under target-disjoint splits, reaching a Pearson r of 0.407 on 90,219 held-out complexes and transferring to 202 independent crystal structures with measured Ki, Kd, IC50 or EC50 at r = 0.467. We report the model against three controls, a target-mean predictor, a ligand-descriptor-only baseline, and within-target correlations, and document both where it performs and where it does not, including its unsuitability for pose rescoring. On an independent 30-compound series against a single GABAA receptor target, PandaDock's empirical scoring function ranks 8th of 25 methods evaluated, ahead of every AutoDock Vina and Vinardo configuration tested, while the GNN scores below Vina, consistent with the within-target ceiling identified on SAIR. At full scale on the PDBbind v2020 refined set (n = 4,640, native crystal poses), the fully independent SAIR model reaches r = 0.531, and a dedicated model trained on PDBbind alone under a target-disjoint split reaches r = 0.690 on its own held-out test complexes, the strongest evidence in this work that PandaDock's affinity predictions generalize. PandaDock is distributed under an open-source license at https://github.com/pritampanda15/PandaDock with a complete command-line interface and a reproducible benchmarking harness.

17
Quantum Encoding Strategies for Drug Response Prediction: An Exhaustive Benchmark on a 20-Qubit Superconducting QPU

Derouich, R.; Mathlouthi, N. E. H.

2026-07-13 bioinformatics 10.64898/2026.07.08.737310 medRxiv
Top 0.1%
1.4%
Show abstract

We present the first systematic, hardware-executed benchmark of twelve distinct quantum data-encoding strategies for drug-response prediction on a real superconducting quantum processing unit (QPU). All experiments were conducted on the IQM Garnet 20-qubit QPU via the IQM Resonance cloud platform, using the Qrisp quantum-software framework (v 0.8.2). Each encoding was evaluated on n = 50 stratified samples drawn from the Genomics of Drug Sensitivity in Cancer dataset (GDSC2, 242 036 drug-cell-line pairs), targeting the natural-log IC50 response variable. Variational weights were optimised offline with the gradient-free COBYLA algorithm before hardware submission. Every circuit was executed with 1024 shots; the regression signal is the zero-qubit Pauli expectation value [&lt;]Z0[&gt;]. Results show that the QAOA-inspired encoding achieves the best RMSE of 3.314 and is statistically superior (p < 0.05, Wilcoxon signed-rank test) to six of the remaining eleven encodings. Hardware-efficient entanglement structures--specifically alternating cost and mixer layers--provide a systematic advantage over purely rotational or diagonal encodings under realistic noise conditions. This work constitutes a reproducible baseline for noise-aware quantum machine learning on pharmaceutical data; all code, data, and raw QPU outputs are publicly released.

18
Sanjeevani: A manually curated anti-cancerous phytochemical database integrated with downstream analysis tools.

Jha, V.; Jha, R.; Shukla, S.; Shingan, S.; Das, G.

2026-06-19 bioinformatics 10.64898/2026.06.15.732344 medRxiv
Top 0.1%
1.1%
Show abstract

BackgroundCancer continues to pose a massive global health burden. While plant-derived phytochemicals offer promising therapeutic leads, existing natural product databases often lack cancer specificity, dataset downloadability, and integrated screening tools. MethodsWe developed Sanjeevani, an integrative web platform cataloguing 4,823 curated anticancer phytochemicals. Using a balanced dataset of 9,646 molecules, we trained Support Vector Machine (SVM), Random Forest, and K-Nearest Neighbours classifiers using a hybrid feature representation of RDKit descriptors and 2048-bit ECFP4 fingerprints. The platform also integrates AutoDock Vina for web-based molecular docking for binding affinity, poses prediction and ADMET-AI for pharmacokinetics estimation. ResultsThe SVM model demonstrated the strongest predictive capability, achieving a top test accuracy of 0.966 and a ROC-AUC of 0.992. Benchmarking across five docking tools confirmed that AutoDock Vina successfully balanced computational automation with literature-consistent binding affinity replication. The final architecture provides rapid interactive 2D/3D visualizations integrated with downstream analysis tools. ConclusionSanjeevani provides an open-access, one-stop pipeline that bridges the gap between raw natural product data and actionable computational screening, accelerating natural product-based oncology drug discovery. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/732344v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@77183borg.highwire.dtl.DTLVardef@d7e465org.highwire.dtl.DTLVardef@1d3dfd7org.highwire.dtl.DTLVardef@10cc94d_HPS_FORMAT_FIGEXP M_FIG C_FIG

19
Muon Reduces the Training Cost of Regulatory DNA Transformers

Doshi, V.; Bhide, M.; Singh, A.; Rathod, Y.; Sivakumar, A.

2026-07-22 genomics 10.64898/2026.07.17.739267 medRxiv
Top 0.2%
1.1%
Show abstract

Gene-therapy design depends on identifying regulatory sequences that drive the right level, timing, and cell-type specificity of expression. Regulatory DNA models offer a way to prioritize such sequences computationally before committing candidates to biological testing. Biological validation involves DNA synthesis, cloning, cell culture, sequencing, and functional screening, so training compute is part of the same constrained discovery pipeline rather than an isolated modeling expense. Reducing the compute required to reach a target pretraining quality could shift time and budget toward larger candidate screens, additional assays, more cell contexts, and broader follow-up validation. Given that Adam-style optimizers are widely used for training genomic sequence models, we study whether Muon can provide a more compute-efficient alternative for regulatory DNA pretraining. We provide an in-depth analysis by training Transformer models (26M-420M parameters) on ENCODE cis-regulatory sequences with Adam and Muon while holding architecture, data, and non-optimizer hyperparameters fixed and varying optimizer family, norm-control scheme, learning rate, and model width. In the largest-scale matched-target comparison, Muon reaches Adam-matched perplexity targets with a median FLOP reduction of 35.4% and a median wall-clock time reduction of 38.5%. The analysis further shows that optimizer rankings depend on norm control: independent weight decay pairs more favorably with Muon than Hyperball in this setting. These findings indicate that optimizer update structure and norm-control choices are practical levers for reducing the training resources required to reach matched perplexity targets in regulatory DNA pretraining.

20
Seeing Nothing, Saying Something: The Lack of Visual Grounding and Confabulation in Gemini Models for Histopathology

Hasan, M. M.; Tozal, M. E.; Ayhan, M. S.

2026-07-07 health informatics 10.64898/2026.07.04.26357257 medRxiv
Top 0.2%
1.1%
Show abstract

Large vision-language models (VLMs) have demonstrated remarkable perfor- mance on computational pathology benchmarks, yet their reliability under adversarial or vacuous inputs remains poorly understood. This paper examines the visual grounding behaviour of two Gemini models Gemini 3.0 Flash Pre- view (gemini-flash) and Gemini 3.1 Pro Preview (gemini-pro) on a well known histopathology classification task, and probes for confabulation using a adver- sarial blank-image set. On the real histopathology dataset both models achieve near-perfect accuracy (98.75% - 100%) across three temperatures (0.0, 0.5, 1.0) and three independent runs. On a controlled adversarial set of blank white images labelled as either benign or malignant, however, a stark divergence emerges. Gemini-flash consistently acknowledges the absence of visual content and assigns zero confidence, while Gemini-pro fabricates detailed, clinically plausible histo- logical descriptions and reports high confidence (mean {approx} 0.95) across the same blank inputs, a behaviour we term confident confabulation. The confabulation rate of gemini-pro reaches 77.8% image-responses at temperature 0.0, dropping to 44.4% at temperature 0.5 and rising to 66.7% at temperature 1.0, while gemini- flash records 0% at all temperatures. These findings raise important questions about the safety and trustworthiness of VLMs in clinical decision-support con- texts, and underscore the need for comprehensive evaluation beyond standard accuracy metrics.